Papers by Juan Pablo Munoz

6 papers
EFTNAS: Searching for Efficient Language Models in First-Order Weight-Reordered Super-Networks (2024.lrec-main)

Copied to clipboard

Challenge: Depending on the size of transformer-based models, they can be restricted from deployment in resource-constrained environments.
Approach: They propose to combine neural architecture search and network pruning techniques to generate and train weight-sharing super-networks that contain efficient transformer-based models.
Outcome: The proposed model achieves high-performing, high-performance subnetworks on the general language understanding evaluation and the Stanford Question Answering Dataset.
LoNAS: Elastic Low-Rank Adapters for Efficient Large Language Models (2024.lrec-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) reach hundreds of billions of parameters and require resources for training and inference stages.
Approach: They propose a low-rank adapter to reduce the number of trainable parameters in a model and reduce memory requirements.
Outcome: The proposed approach reduces memory and compute requirements while preserving performance.
Mamba-Shedder: Post-Transformer Compression for Efficient Selective Structured State Space Models (2025.naacl-long)

Copied to clipboard

Challenge: Large pre-trained models have achieved outstanding results in sequence modeling . alternative architectures, such as Selective Structured State Space Models (SSMs), have been proposed to address these inefficiencies.
Approach: They propose to reduce the size and computational overhead of large pre-trained models by removing selected components at different granularities.
Outcome: The proposed models achieve a speedup of up to 1.4x during inference while maintaining accuracy.
RTTC: Reward-Guided Collaborative Test-Time Compute (2025.findings-emnlp)

Copied to clipboard

Challenge: Reward-Guided Test-Time Compute (RTTC) is a powerful paradigm for large language models . indiscriminate application of TTC strategy incurs substantial computational overhead .
Approach: They propose a framework that adaptively selects the most effective TTC strategy for each query via a pretrained reward model.
Outcome: The proposed framework maximizes accuracy across diverse domains and tasks.
FedReFT: Federated Representation Fine-Tuning with All-But-Me Aggregation (2026.findings-eacl)

Copied to clipboard

Challenge: Representation Fine-Tuning (ReFT) adapts large pre-trained models by updating only a small subset of parameters.
Approach: They propose a method that uses sparse intervention layers to steer hidden representations directly to capture rich semantic information.
Outcome: The proposed approach outperforms PEFTs on commonsense reasoning, arithmetic reasoning, and GLUE benchmarks while maintaining a high parameter efficiency.
Federated Foundation Models: Privacy-Preserving and Collaborative Learning for Large Models (2024.lrec-main)

Copied to clipboard

Challenge: Foundation Models (FMs) have demonstrated success in a wide range of applications, but their optimization often requires access to sensitive data.
Approach: They propose a framework that combines FMs and Federated Learning to enable privacy-preserving and collaborative learning across multiple end-users.
Outcome: The proposed framework combines benefits of FMs and Federated Learning (FL) it enables privacy-preserving and collaborative learning across multiple end-users.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations